Papers with spatial reasoning capabilities

4 papers
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners (2024.emnlp-main)

Copied to clipboard

Challenge: Top-view perspective is a typical way in which humans read and reason over different types of maps, but spatial reasoning capabilities of modern VLMs in this setup remain unattested and underexplored.
Approach: They introduce a top-view spatial reasoning dataset and use it to evaluate VLMs across 4 perception and reasoning tasks with different levels of complexity.
Outcome: The proposed model can understand and reason over spatial relations from the top view and can be controlled at different granularities of spatial reasoning.
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks explore aspects of threedimensional spatial reasoning and visual-language reasoning in dynamic environments, but they are unable to perform well on 3D spatial deformation reasoning.
Approach: They propose to use a ladder competition format to assess the model's spatial deformation reasoning abilities to determine its performance.
Outcome: The proposed framework assesses the performance of Vision-Language Models in spatial deformation reasoning tasks.
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) have shown impressive generalization ability on vision and language tasks, but their spatial understanding is under-explored.
Approach: They construct a VQA dataset to analyze LMMs' spatial reasoning capabilities.
Outcome: The proposed model is stronger at basic object detection than complex spatial reasoning.
iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) struggle with spatial reasoning and visual alignment, despite their performance on 2D tasks.
Approach: They propose a multimodal benchmark to evaluate VLMs' spatial reasoning capabilities based on the sliding tile puzzle .
Outcome: The proposed model performs better on 2D tasks compared to 3D or text-based settings, but struggles with complex spatial configurations and consistently falls short of human performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations